Research Synthesis Methods
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match Research Synthesis Methods's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Dobin, D.; Witmer, A. M.; Sweeney, F.; Ryan, T.; Cimino, A.; Haroz, E. E.; Nestadt, P. S.; Wilcox, H. C.
Show abstract
Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.
LIn, H.; Lyu, J.
Show abstract
BackgroundQuality Control Circle (QCC) reports are often reviewed qualitatively, but reviewer workload and inter-rater variability make large-scale assessment difficult. We evaluated whether multiple large language models (LLMs) could score QCC methodological quality reliably on a designed-anchor benchmark. ObjectiveTo estimate inter-model reliability for QCC quality scoring and to assess whether model scores align with designed synthetic anchors and remain descriptively comparable to a small set of public PMC QCC reports. MethodsWe evaluated 30 synthetic QCC reports and 8 public PMC QCC reports across four primary evaluators (GPT, Gemini, Grok, DeepSeek) and one sensitivity evaluator (Claude); Claude was excluded from the primary panel because it shared the model family used during prompt development. Each synthetic case was scored across eight QCC quality dimensions in three runs per evaluator. We summarized each evaluator by median scores, then estimated ICC(A,1) across the primary panel. We also examined score-based calibration against designed anchors, keyword-assisted defect mention, leave-one-out and k=5 sensitivity, and a descriptive synthetic-versus-PMC distributional plausibility check. ResultsInter-model reliability on the primary k=4 panel was excellent: ICC(A,1) = 0.953 (95% CI 0.944 to 0.962) with 237 pooled case-dimension rows. The pre-specified k=5 sensitivity analysis including Claude was 0.954, and leave-one-out estimates within the primary panel ranged from 0.950 to 0.959. Score-based calibration against designed anchors met the prespecified target in 57/58 trap-affected case-dimension rows (98.3%). Keyword-assisted defect mention was present in 51/58 trap instances (87.9%). The synthetic-versus-PMC comparison was descriptively similar across all eight dimensions, and all dimensions met the predefined descriptive margin check. ConclusionsIn this designed-anchor pilot, multi-model LLM scoring of QCC methodological quality showed high inter-model reliability and stable alignment with synthetic anchor scores. These findings support benchmark feasibility, but they do not establish expert validity, clinical validity, or operational deployment readiness.
Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.
Show abstract
Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.
Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.
Show abstract
Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.
Moe-Byrne, T.; Knapp, P.; Golder, S.
Show abstract
Background People with lower levels of literacy or health literacy may struggle to understand conventional health information. Video animations show promise as information tools, yet it is unclear whether video animations help reduce these inequalities in understanding. This study examined whether the effectiveness of video animations in health settings differs according to level of literacy or health literacy. Methods We drew on trials from a recent systematic review of video animations about healthcare or public health topics for patients or the public. We extracted available data on literacy, health literacy, or proxy indicators. One reviewer extracted data and a second checked all entries. Where possible, we conducted subgroup analyses of low and high literacy levels or interaction meta-analyses comparing low versus high literacy groups; otherwise, results were summarised narratively. Results From 88 eligible trials, we extracted health literacy data for 12. Across nine trials reporting knowledge, animations mostly improved knowledge compared with controls in both lower and higher health literacy groups. Effects on attitudes and behaviours were mixed and often small, with few studies reporting results by health literacy level. Across the subgroup analyses available, there was no consistent evidence of a pooled interaction effect of animations according to low and high literacy groups, but both statistical heterogeneity and small subgroup sizes limited precision of estimates. Across 88 trials, 54 (61%) reported education level, 22 (25%) did not, and 12 (14%) involved children or adolescents likely to have similar education levels. Conclusions Overall, the available data suggest that video animations can improve knowledge outcomes in both lower and higher health literacy groups, but their impact on attitudes and behaviour is less clear. Because literacy was rarely reported or analysed in the trials, it remains uncertain whether animations help to reduce literacy-related inequalities in access to, and use of health information.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Bunning, B. J.; Weng, Y.; Wu, D. J.; Hui, G.; Hope, J. E.; Pandurangan, V.; Lopez, I.; Everett, S.; Chen, J. H.; Desai, M.
Show abstract
Doctors increasingly rely on AI in the clinic, yet which report features make AI-generated responses useful and trustworthy remains unclear. In this randomized mixed-methods study, 34 oncology physicians provided 294 ratings of four blinded AI systems across five vignettes, alongside 20 semi-structured interviews analyzed with a prespecified LLM-assisted qualitative pipeline. Despite similar references, an evidence-graded report adapted from OpenEvidence was rated significantly lower in overall utility than standard OpenEvidence (mean difference, -0.96; 95% CI, -1.26 to -0.66; P<.001). Qualitative analysis identified six themes and seven design requirements. Oncologists valued rapid orientation, evidence retrieval, and verification, preferring concise, scannable reports with quantitative outcomes, recognizable bolded guidelines, explicit uncertainty, and verifiable citations. Trust deteriorated with citation mismatch, buried provenance, evidence misclassification, overconfident recommendations, and poor organization. Evidence presented differently can alter perceptions of clinical utility and trust; accuracy alone is insufficient, and report design must also be empirically evaluated.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Lemarchand, C.; Naudet, F.; Pencole, M.-A.; Ropers, L.; Scanff, A.; Cristea, I. A.; Locher, C.
Show abstract
Objective: In 2021, a large-scale survey highlighted that in a subset of biomedical journals, a few authors -often serving on the editorial board- published disproportionately and experienced shorter acceptance times. Our study aims to specifically quantify editors research articles within the journals in which they operate. Methods: We selected journals indexed in Open Editors, a dataset that collects publicly available information on journal editorial boards through web scraping. Journals not indexed in PubMed, mega-journals, and those with very low publication volume were excluded. For the remaining journals, we linked the 2022 editorial boards from Open Editors to authors of research articles (i.e., original articles, case reports, and reviews) published between 2020 and 2023. For each journal, we then computed indicators describing publication patterns: the percentage of research articles (i) by the most prolific editor, (ii) with at least one editor, and (iii) by the most prolific author, as well as publication lags for each article. Results: Across the 1,623 journals studied, the median and 95th percentile of research articles are 1.78% and 6.7% for those co-authored with the most prolific editor, 12.1% and 41.1% for those with at least one editor, and 2.5% and 7.9% for those with the most prolific author. An editor was among the most prolific author(s) in 45.0% of the journals. For authors, the median and 5th percentile publications lags are 99 and 35 days; for editors, it is 95 and 33 days; and for editors-in-chief, it amounts to only 84 and 12 days. An in-depth examination of journals where the most prolific editor co-authored more than 6.7% (95th percentile) found a median impact factor of 3, and a median h-index of 42 for their most prolific editor(s). Conclusion: In 5% of cases, an editor contributes to approximately >7% of the articles published in their own journal. In nearly half of the journals, the most prolific author is an editor. These results need to be complemented by a qualitative approach to examine whether research articles authored by editors appropriately address potential conflicts of interest, as required by COPE recommendations, and to better understand the motivations underlying this practice.
Bergman, H. I.; Liu, V.; Austin, B.; Ali, S.; Fiedler, M.; Sandiford, C.; Blanchard, R.; Casanovas, C. L.; Pedrazzini, G.; Markopouliotis, T.; Vermersch, F.
Show abstract
Background Ambient AI documentation tools, known as scribes, are entering routine clinical practice at scale, but the evidence comparing the notes they produce against clinician-written notes is dominated by single-site, single-language studies that rely on human review to find errors, a method known to miss most documentation errors. Methods We conducted a paired simulation across five countries and languages (Cambridge/English, Barcelona/Spanish, Milan/Italian, Paris/French, Cologne/German; 385 paired consultations, 770 notes). From each actor-performed consultation, an AI scribe (Heidi) and a junior-to-middle-grade clinician independently produced a note. Notes were scored on the PDQI-9 by evaluators blinded to authorship. Documentation errors were identified by two methods of deliberately different sensitivity - clinician adjudication, and a calibrated automated reviewer externally validated against a blinded ten-clinician panel - then graded for clinical risk by a three-model panel. The co-primary outcomes were PDQI-9 total and Critical+High error burden, the latter reported under both detection arms. The analysis plan was registered before any pooling across sites. Results AI notes scored higher than clinician notes on the PDQI-9 (40.6 vs 35.6; difference +5.08, 95% CI 4.6-5.6; Cohen dz=0.55), consistently across all five sites (dz 0.41-0.75), and were less dispersed (5.7% of AI vs 27.8% of clinician notes fell below the study pre-specified low-score threshold (<32)). On the principal safety outcome - the paired probability that a note carried [≥]Critical+High error - clinician notes were affected more often under both detection arms: 61.0% versus 24.4% by the calibrated reviewer (relative risk 2.50, 95% CI 2.09-3.00) and 21.8% versus 6.2% by clinician adjudication (relative risk 3.50, 95% CI 2.32-5.27). The difference was largest for omissions. Unaided clinician review identified roughly 12% of the errors the calibrated reviewer retained, and a smaller fraction in AI notes than in clinician notes. Conclusions In this simulation, AI-generated notes scored higher on documentation quality, varied less, and carried fewer clinically significant errors than notes written on the same consultations by junior-to-middle-grade clinicians. The magnitude of the safety difference depends on the sensitivity of error detection, so we report both detection regimes and bound rather than point-estimate the absolute error rate. Extension to live practice, consultant-authored documentation, and notes as filed after clinician editing remains to be established.
Delporte, M.; Tamimi, R.; Mehta, S.; Choi, E.; Zhang, Y.; Shi, Y.
Show abstract
Objective To develop and evaluate an automated large language model (LLM)-based framework for conducting meta-analyses of nutrition-related exposures and the risk of breast, ovarian, and uterine cancers. Design We developed MetaFemina, an automated evidence-synthesis pipeline for women's cancers that integrates keyword-based literature retrieval, LLM-assisted evidence extraction, and random-effects meta-analysis. We evaluated its performance against two recently published peer-reviewed meta-analyses and compared exposure-outcome associations across the three cancer types. Data sources PubMed articles identified through keyword-based searches of titles and abstracts. Methods MetaFemina was developed as a web platform that identifies relevant scientific articles, automatically extracts relevant information using LLMs, and synthesizes extracted evidence using random-effects meta-analysis. Additional analyses included assessment of heterogeneity, publication bias, and leave-one-out sensitivity analyses. The platform also provides sample size calculations based on synthesized effect sizes and generates visual summaries and plain-language interpretations. Results Compared with two recent peer-reviewed meta-analyses of folate and vitamin E intake in relation to breast cancer risk, MetaFemina demonstrated high sensitivity (81.82% and 80%, respectively) in identifying eligible studies and additionally retrieved relevant articles that had been missed by manual screening (27 and 13, respectively). Among 226 exposures considered, lutein and beta-carotene were significantly associated with lower risks of breast, ovarian, and uterine cancers. Vitamin D, antioxidants, and soy were significantly associated with lower risks of both breast and ovarian cancers, whereas calcium and folic acid were significantly associated with lower risks of both breast and uterine cancers. In contrast, iron, red meat, and copper were significantly associated with higher risks of both breast and uterine cancers. omega-6 fatty acids showed contrasting associations, being significantly associated with higher breast cancer risk but lower ovarian cancer risk. After restriction to dietary-intake studies, these cross-cancer significant associations remained statistically significant except for copper, which no longer met the two-study threshold for either breast or uterine cancer. Additionally, calcium became significantly associated with lower ovarian cancer risk, resulting in significant negative associations across all three cancer types, while vitamin E became significantly associated with lower breast cancer risk and remained significantly associated with lower ovarian cancer risk. Conclusions MetaFemina demonstrated high sensitivity for identifying relevant scientific literature, extracts key evidence, and performs statistically rigorous automated meta-analyses. The framework may facilitate more rapid evidence synthesis in nutritional epidemiology and may support researchers in study design, hypothesis generation, and interpretation of emerging evidence.
Maleki, C.; Bertrand, Y.; Gailly, F.
Show abstract
Clinical recommendations are often expressed in narrative form, which limits their direct execution, auditability, and patient-specific interpretation. This paper presents a hybrid decision-support framework that combines Decision Model and Notation (DMN), survey-weighted rule-ensemble learning, and counterfactual sensitivity analysis. The framework is evaluated using an NHANES-derived fasting cohort for classification of documented diabetes status. The full fasting analysis cohort contained 2,582 participants, and a non-diagnostic laboratory subgroup, Gate0, contained 2,111 participants. On untouched test data, the rule-ensemble model achieved ROC-AUC and PR-AUC values of 0.959 and 0.873 in the full fasting cohort and 0.861 and 0.499 in Gate0. Four clinically interpretable candidate rules were selected using validation data only. A nonnegative survey-weighted logistic model removed one redundant rule and converted the remaining three binary activations into an auditable DMN score and model-estimated probability. The final DMN achieved ROC-AUC 0.769, PR-AUC 0.153, and Brier score 0.029 in the untouched Gate0 test set. In small rule-defined test subgroups, hypothetical five-unit BMI reductions lowered mean model-estimated probability by 2.40 to 5.89 percentage points when one or more BMI thresholds were crossed. These findings characterize policy sensitivity rather than causal effects and require external validation.
Krump, P. A.; Blasingame, M. N.; Koonce, T. Y.; Williams, A. M.; Su, J.; Giuse, N. B.
Show abstract
Background: Large language models (LLMs) that use retrieval-augmented generation (RAG) are increasingly used to answer clinical questions, although the evaluation of these systems remains limited. Building on previous studies conducted by our team, this case report aimed to improve upon this knowledge gap by applying a reusable methodology to compare the performance of eight LLMs that utilize RAG techniques for evidence synthesis. Case Presentation: Eight commercially available RAG LLM tools (OpenEvidence, Undermind, Consensus, SciSpace, Elicit, MediSearch, EvidenceHunt, and Scite) were evaluated using twelve ChatGPT-generated clinical questions on the topics of treatment, etiology, and prognosis. To enable comparison, we prompted ChatGPT to identify all key unique medical concepts from the full set of LLM responses to each question. Concepts were categorized as critical ("must-have") or non-critical ("nice-to-have") for answering the clinical question. Experienced information scientists were consulted at each step for their expertise. Descriptive statistics and Kruskal-Wallis tests were used to compare performance across tools and question categories. No significant differences were found among the eight RAG LLMs in their coverage of "must-have" (p=0.95) or "nice-to-have" (p=0.16) key unique medical concepts, and no single tool consistently captured all identified concepts. Conclusions: These findings suggest that RAG LLMs may be supplementary tools for evidence retrieval and synthesis but cannot, at this time, fully replace comprehensive expert review of the medical literature. The evaluation framework presented here may be a useful model for future comparative assessments of rapidly evolving AI evidence synthesis tools.
Austria, D.; McCollister, B.; Lindsey, J. E.; Arowolo, M.; Okon, M.
Show abstract
Objective. Formal large language model (LLM) evaluations score isolated prompts, but clinicians and health-informatics researchers meet model failures inside multi-step workflows where erroneous output can alter procedures or contaminate documents. We present TRACE (Tracking Reliability of AI-generated Conversational Evidence), a practitioner-audit framework for evaluating the downstream workflow reliability of conversational AI. Materials and Methods. A method paper with an empirical demonstration: 45 documentation-positive incidents recorded by one clinician-informatician across scholarly, clinical informatics, and clinical-adjacent workflows over seven weeks, coded with a consequence-based severity rubric, an error definition, a taxonomy crosswalk, and a Response-Audit Scorecard. Three reviewer-authors independently coded a 16-incident subsample; three vendor-blinded AI comparators applied the taxonomy to all 45 incidents. Results. Four categories tied as most frequent: verification failure, factual numerical error, tool-behavior misunderstanding, and citation or reference formatting (n=7 each). Four workflow-harm patterns recurred: procedural propagation, documentary contamination, trust-calibration disruption, and user-borne corrective burden, and one incident carried an estimated $2500 impact. Category agreement across three human reviewer-authors was low (Fleiss {kappa}=0.155), whereas three AI comparators agreed substantially (Fleiss {kappa}=0.632), suggesting taxonomy legibility under standardized conditions even where human judgment diverged. Discussion. Category assignment is comparatively legible, whereas severity and claimed-verification remain judgment-dependent. The claimed-verification gap is a measurable failure mode distinct from hallucination, sycophancy, and over-refusal. Conclusion. Practitioner audits with structured response scoring complement benchmarks by documenting workflow harm as an applied evaluation unit for clinical informatics and public-health work; this is a pilot that motivates, not estimates, error rates or cross-model comparisons.
Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.
Show abstract
We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.
Chen, Y.; Popescu, M.
Show abstract
Background: Clinical terminology pipelines must first extract candidate spans from narrative notes and then determine whether those spans map to existing concepts or warrant further review. Evaluation is difficult because span boundaries vary between annotators and because downstream decisions depend on the terminology evidence retrieved for each span. Objective: We evaluated clinical concept extraction, terminology linking across controlled evidence conditions, and ontology-extension triage for terms that remained unmatched after initial terminology screening. Methods: We conducted 3 complementary pilot evaluations that used distinct units of analysis and were analyzed separately. Study 1 compared 5 automated extraction pipelines and a union-merge analysis with 2 human annotation sets in 66 deidentified clinical notes from 3 health systems. Agreement was evaluated by exact string matching and BGE-large-en-v1.5 embedding matching. Study 2 evaluated 56 clinical spans, including 28 with reference Unified Medical Language System concepts and 28 adjudicated as unsuitable for ontology extension, under complete retrieval, matched-concept masking, and large language model-only inference, yielding 168 span-condition outputs. The graph retrieval pipeline used BGE-large-en-v1.5 embeddings, and the decision model was Gemma 3 27B. Study 3 applied full vector retrieval to 84 terms previously not matched in either UMLS or BioPortal. Results: In Study 1, interannotator exact-match F1 was 0.29 and embedding-match F1 was 0.75. Automated exact-match F1 scores ranged from 0.07 to 0.17; embedding-match F1 was highest for MedGemma (0.55), followed by Gemma (0.53), sci_md and SciBERT (each 0.43), and Llama 3.3 (0.32). In Study 2, complete retrieval returned a reference-matched link for 28/28 known-concept spans (100%; 95% CI, 87.9%-100%). Masking assigned POSSIBLE_CANDIDATES to all 28; large language model-only inference assigned POSSIBLE_CANDIDATES to 25/28 (89.3%) and LINKED to 3/28 (10.7%). Across the 3 evidence conditions, the same 12/28 unsuitable-extension spans were classified as NOT_MEANINGFUL (42.9%) and the same 16/28 as POSSIBLE_CANDIDATES (57.1%). In Study 3, the pipeline assigned PLAUSIBLE_EXISTING_CONCEPT to all 84 terms, none was flagged for extension, and top-candidate similarity averaged 0.914 (SD 0.027); extension status was not independently adjudicated. Conclusions: Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence. In the follow-up sample, initial nonmatching did not establish ontology novelty: after semantic retrieval, the pipeline classified all 84 terms as plausible existing concepts and proposed none for extension. These findings support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.
MUTHUKA, J. K.; Nyambura, L. W.; Onyango, C. K.; Oluoch, K.; Kioko, M.; Maluki, J.; Nzioki, J. M.; Kim, S.
Show abstract
Background: Autism spectrum disorder (ASD) is a lifelong neurodevelopmental condition for which timely diagnosis is critical to early intervention, family support, and equitable access to care. However, substantial disparities in access to ASD diagnostic services persist across socioeconomic, geographic, clinical, and health-system contexts. This systematic review and meta-analysis synthesized evidence on determinants of access across the ASD diagnostic pathway, from recognition and referral to diagnostic completion and timely diagnosis. Methods: We systematically searched MEDLINE/PubMed, Embase, Scopus, Web of Science, Global Health, and grey-literature sources for studies published between January 2004 and December 2024. Eligible studies examined determinants of ASD diagnostic completion, diagnostic pathways, diagnostic timeliness, or barriers and facilitators to diagnostic access. Two reviewers independently extracted data and assessed methodological quality using the Mixed Methods Appraisal Tool (MMAT). Quantitatively comparable estimates were synthesized using random-effects models with restricted maximum likelihood estimation. Heterogeneity was assessed using Cochran's Q, I2, tau2, and 95% prediction intervals. Pre-specified subgroup analyses, meta-regression, sensitivity analyses, funnel-plot assessments, and Bayesian random-effects analyses were undertaken. Results: The search identified 4,899 records; after removal of 537 records without associated data, 4,362 records underwent title/abstract screening. 3,800 records were excluded, 562 reports were sought for retrieval, and 450 full-text reports were assessed after 112 could not be retrieved. Ultimately, 22 unique studies met the inclusion criteria. Nine unique studies contributed 23 quantitative effect estimates, while the remaining studies contributed to the narrative synthesis. The evidence covered socioeconomic, geographic, family, communication, screening, child developmental, provider, and health-system determinants. The overall random-effects meta-analysis yielded a pooled diagnostic access outcome of 74.1% (95% CI 65.8-81.1%), with substantial heterogeneity (Qe=209.95, p<0.001; I2=88.4%, 95% CI 79.1-94.4%; tau2=0.691) and a wide 95% prediction interval of 32.8-94.4%. Bayesian analysis produced a highly concordant pooled estimate of 73.3% (95% CrI 65.3-80.2%), with I2=87.5% and tau=0.833, and satisfactory MCMC convergence (R-hat=1.000). By outcome domain, pooled successful outcomes were highest for diagnostic pathways (89.3%, 95% CI 70.1-96.7%), followed by timely diagnosis (76.3%, 95% CI 62.9-86.0%), and lowest for diagnostic completion (67.1%, 95% CI 61.8-72.0%) (Qm=5.98, p=0.050). Timely diagnosis demonstrated particularly high heterogeneity (I2=91.2%), whereas diagnostic completion showed moderate heterogeneity (I2=40.6%). Across determinant domains, frequentist pooled estimates were 79.5% for child developmental/neurobehavioral factors, 74.2% for family/socioeconomic/perceptual factors, 68.0% for intervention/care-navigation factors, and 63.6% for provider/clinical recognition factors. Bayesian estimates were 76.7% (BF=53.76), 72.9% (BF=226.32), 64.3% (BF=25.60), and 53.7% (BF=0.684), respectively. Meta-regression indicated that determinant category (Qm=13.48, p=0.004) and effect measure (Qm=7.81, p=0.020) significantly explained between-study variation, whereas age group (p=0.203) and geographic region (p=0.453) did not. Family/socioeconomic factors had significantly larger effect sizes (B=2.703, 95% CI 0.661-4.744; p=0.009), as did child developmental/neurobehavioral factors (B=1.516, 95% CI 0.047-2.985; p=0.043). Potential small-study effects were detected by two of three asymmetry tests, although the Rosenthal fail-safe N was 1,723. Trim-and-fill identified seven potentially missing estimates, with an adjusted pooled effect of 68.4% (95% CI 27.7-109.1%). Importantly, exclusion of two influential outlying estimates produced a pooled outcome of 77.1% (95% CI 71.6-81.9%), indicating that the principal finding was robust. Conclusions: Approximately three-quarters of observed ASD diagnostic outcomes represented successful access, but the substantial heterogeneity indicates that diagnostic access is highly context-dependent. Families were more likely to successfully navigate diagnostic pathways than to complete diagnostic assessment, while timely diagnosis showed the greatest variability across settings. Family and socioeconomic circumstances and child developmental characteristics emerged as particularly important determinants, whereas provider-related effects were more heterogeneous and uncertain. Improving equitable ASD diagnosis requires interventions spanning the entire diagnostic pathway, including developmental surveillance, screening, referral coordination, family navigation, provider capacity, specialist availability, and mechanisms to ensure completion of diagnostic assessment. Greater longitudinal and implementation research is particularly needed in low- and middle-income countries, where diagnostic infrastructure and specialist capacity remain limited.
Ndabashinze, R.; Franzen, D.; Kozuch, E.; Aagerup, J.; Fink, A.; Yerunkar, S. S.; Hunter, K.; Mayo-Wilson, E.; Ying, X.; Kilicoglu, H.; Schorr, S. G.; Seidler, A. L.
Show abstract
Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.
Merlo, J.; Bashir, N. Z.; Rodriguez-Lopez, M.; Khalaf, K.; Öberg, J.; Perez-Vicente, R.
Show abstract
Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) describes health inequalities through three components: (i) specific contextual effects (SCE), (ii) general contextual effects (GCE), and (iii) discriminatory accuracy of the context. We present Simple-Means MAIHDA (S-MAIHDA), which estimates each stratum directly from its observed individuals, with no distributional assumption. The observed proportions are unbiased whatever the stratum size, and their confidence intervals report the uncertainty honestly. S-MAIHDA operationalises the three components on the probability scale. The SCE are the raw and standardised stratum prevalences and the modification of the sociodemographic average differences by the area. The GCE are the variance partition coefficient (VPC) and the contextual structuring of the between-stratum inequality, expressed as the contextual clustering of inequalities, the additive sociodemographic differences, and the contextual modification of inequalities (CMI). The contextual discriminatory accuracy is expressed by the area under the ROC curve (AUC), and the sensitivity and specificity at the population prevalence as the threshold for a possible intervention. Because its estimates are the observed data themselves, S-MAIHDA is the canonical description, and the compare diagnostic quantifies how Random-Effects MAIHDA (RE-MAIHDA), the usual implementation, departs from it: RE shrinkage pulls small strata towards the overall mean and can hide the very inequalities the analysis seeks. The approach is implemented in the smaihda Stata command and reproduced in free Python code. We illustrate S-MAIHDA on register data from Malmo, Sweden (43,291 individuals; 300 area-sociodemographic strata), showing how the three components separate two contrasting outcomes: psychotropic medication use, almost purely sociodemographic, stable across areas, with weak contextual structuring (VPC {approx} 4%, CMI {approx} 0%); and choice of a private general practitioner, strongly geographical (VPC {approx} 11%, CMI {approx} 17%), with the sociodemographic differences reshaped and amplified in wealthy areas. RE-MAIHDA attenuated inequalities. For describing inequalities, S-MAIHDA preserves what the data show.
Liu, A.; Ho, A.; Droste, A. M.; Martin, D.; Wong, E.; Zhou, E.; Zhou, I.; Park, J.; Jiao, J.; Skelly, K.-R.; Kim, K.; Li, J.; Rao, K.; Uehara, M.; Marion, M.; Fitzgerald, N.; Dias, R.; Shringarpure, S.; Yuan, Y.; Wang, Y.
Show abstract
We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. The majority of existing life sciences benchmarks have a narrow scope or are purely knowledge-based, and therefore fail to capture the complexity of real-world research, which often involves ambiguities and requires the accurate execution of multiple dependent judgment calls. Additionally, almost all existing benchmarks span at best a small collection of subdomains within the life sciences; there is at present no existing life sciences benchmark with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional research settings. LifeSciBench addresses this gap by spanning seven representative scientific workflows and seven life science domains, with each constituent task paired with a human expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best, with a task-weighted mean normalized rubric score of 0.576 and a task-weighted response pass rate of 36.1% (response-level values are first averaged within each task, and the resulting task-level values are then averaged with equal weight). LifeSciBench remains unsaturated, with 171 tasks (22.8%) having no observed passing response from any evaluated model and 261 tasks (34.8%) having a best-model pass rate below 20%. LifeSciBench therefore serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in the life sciences.